Papers by Kush R. Varshney
Evaluating the Prompt Steerability of Large Language Models (2025.naacl-long)
Copied to clipboard
Erik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy, Elizabeth M. Daly, Kush R. Varshney, Eitan Farchi, Pierre Dognin, Jesus Rios, Djallel Bouneffouf, Miao Liu, Prasanna Sattigeri
| Challenge: | a primary question underlying alignment research is: whose views are we aligning to? |
| Approach: | They propose to evaluate the steerability of model personas as a function of prompting by defining a benchmark and inspecting how these indices change as if steering effort is a factor. |
| Outcome: | The proposed benchmark reveals that the steerability of many current models is limited due to skew in baseline behavior and an asymmetry in their steerability across many persona dimensions. |
Granite Guardian: Comprehensive LLM Safeguarding (2025.naacl-industry)
Copied to clipboard
Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia, Subhajit Chaudhury, Tejaswini Pedapati, Pierre Dognin, Keerthiram Murugesan, Erik Miehling, Martín Santillán Cooper, Kieran Fraser, Giulio Zizzo, Muhammad Zaid Hameed, Mark Purcell, Michael Desmond, Qian Pan, Inge Vejsbjerg, Elizabeth M. Daly, Michael Hind, Werner Geyer, Ambrish Rawat, Kush R. Varshney, Prasanna Sattigeri
| Challenge: | a suite of advanced models is designed to detect and mitigate risks associated with prompts and responses. |
| Approach: | a team of researchers develop a model family to detect and mitigate risks associated with prompts and responses. the model family is based on the Granite 3.0 language models. |
| Outcome: | a new model family is designed to detect and mitigate risks associated with prompts and responses. |
AI Steerability 360: A Toolkit for Steering Large Language Models (2026.acl-demo)
Copied to clipboard
Erik Miehling, Karthikeyan Natesan Ramamurthy, Praveen Venkateswaran, Ching-Yun Ko, Pierre Dognin, Moninder Singh, Tejaswini Pedapati, Avinash Balakrishnan, Matthew Riemer, Dennis Wei, Inge Vejsbjerg, Elizabeth M. Daly, Kush R. Varshney
| Challenge: | The AI Steerability 360 toolkit is an extensible, open-source Python library for steering LLMs. |
| Approach: | The AI Steerability 360 toolkit is an extensible, open-source Python library for steering LLMs. |
| Outcome: | The toolkit is available under an Apache 2.0 license and is available on https://github.com/IBM/AISteer360. |